perf: make quoted-printable decoding beat 0.9.0 on dense escapes too - #264
Merged
Merged
Conversation
Comparing the 0.9.0 release against master turned up exactly one benchmark where the release still won: parse_qp_dense_escapes, by 58% on the x86 gate and 20% on an M4. #229 replaced the quoted_printable crate with a run-copying decoder and measured -39%. That was true of the benchmarks that existed. The dense shape did not have one: parse_qp_dense_escapes arrived later with #223, so the one input the new decoder is worst at was never compared against the code it replaced. Measured directly, ours took 236 us against the crate's 166 us. Two changes, both in the vendored qp.rs: decode_line now looks at the byte under the cursor before calling memchr. After an escape the next byte is very often another `=` -- every non-ASCII character encodes as two or three consecutive escapes -- so memchr was being called only to report a match at offset 0, paying its SIMD setup each time. On the dense fixture that was 40,000 calls. The rule-1 pre-scan finds the first dropped byte a chunk at a time rather than with `position`, which cannot vectorise because it must stop at the first hit. `is_kept` is respelled as arithmetic so the chunk reduction is branchless; is_kept_is_the_same_set checks all 256 bytes against the original spelling. On 120 KB with nothing to drop -- nearly every real body -- that scan goes from 63 us to 5.5 us. Local interleaved A/B, M4, 3 rounds, controls within 3.1%: parse_qp_dense_escapes 0.171 -> 0.102 ms (-40%), parse_qp_message 0.129 -> 0.085 ms (-34%), everything else inside the floor. Against 0.9.0 the dense case is now 1.40x faster instead of 1.20x slower, and ordinary quoted-printable 2.6x faster. Output is unchanged. The vendored crate-agreement tests pass, including all 65,536 `=xy` byte pairs, and qp_agreement fuzzed 7.3M executions against the crate with no disagreement. upstream.patch is regenerated and PATCH.md records both changes and why they are shaped that way. Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Comparing the 0.9.0 release against master turned up exactly one benchmark where the release still won:
parse_qp_dense_escapes, by 58% on the x86 gate and 20% on an M4. This closes that gap and then some.Why it was there
#229 replaced the
quoted_printablecrate with a run-copying decoder and measured −39%. That was true of the benchmarks that existed — butparse_qp_dense_escapesarrived later, with #223, so the one input the new decoder is worst at was never compared against the code it replaced.Measured directly on
=C3=A9× 20000 (120 KB, an escape every three bytes):The two changes
Both in
vendor/mailparse/src/qp.rs.1. Look under the cursor before reaching for
memchr. After an escape the next byte is very often another=— every non-ASCII character encodes as two or three consecutive escapes — somemchrwas called only to report a match at offset 0, paying its SIMD setup each time. On the dense fixture that was 40,000 calls. This is why the fix helps ordinary mail too, not just the pathological case.2. Find the first dropped byte a chunk at a time. The rule-1 pre-scan used
position, which cannot vectorise because it must stop at the first hit. It now reduces 32-byte chunks with|=and only falls back to a byte-wise scan inside the chunk that failed.is_keptis respelled as arithmetic (b - 0x20 < 0x5F,b - 9 < 2,b == 0x0D) so that reduction is branchless. On 120 KB with nothing to drop — nearly every real body — 63 µs → 5.5 µs.Results
Local interleaved A/B, Apple M4, 3 rounds, controls within 3.1%:
parse_qp_dense_escapesparse_qp_messageEverything else is inside the noise floor. Against 0.9.0 the dense case is now 1.40× faster instead of 1.20× slower, and ordinary quoted-printable 2.6× faster.
Output is unchanged
is_keptwas respelled, so the risk is that the set changed.is_kept_is_the_same_setchecks all 256 bytes against the originalmatches!spelling. On top of that the existing crate-agreement tests still pass — the generated corpus at every length, all 65,536=xybyte pairs, the per-rule cases and the real fixture — andqp_agreementfuzzed 7.3M executions against the crate with no disagreement.upstream.patchis regenerated (check_vendored_mailparse.shpasses) andPATCH.mdrecords both changes and why they are shaped that way, so the next mailparse hand-merge does not quietly drop them.